Skip to content

test: validate current vLLM cache-source metrics on DeepSeek-V4-Pro - #3491

Draft
cquil11 wants to merge 9 commits into
mainfrom
codex/pr56318-bd57138-agentx
Draft

cquil11 wants to merge 9 commits into
mainfrom
codex/pr56318-bd57138-agentx

Conversation

@cquil11

@cquil11 cquil11 commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Rerun vLLM #56318 at 6ce00ee2ebcf29076d671786c36bf8f8ca71bd0b on DeepSeek-V4-Pro.

  • B200 TP8: native DRAM c8, Simple NVMe c14, native DRAM + NVMe c14.
  • GB300: NIXL + Mooncake Store c256, one DEP4 prefill worker and one DEP16 decode worker. Five nodes, 20 GPUs.
  • AIPerf AgentX, 3,600 seconds per point, MTP3, no evaluations, Python frontend.
  • Preserve the proven serving settings and runtimes. Refresh the attribution overlay with bounded attention ranges and sweep-line accounting.
  • Docker Hub images pinned by digest. Workers verify overlay checksums before model startup. AIPerf collects Prometheus source counters, including all five GB300 DP endpoints.

Previous revision: all four points passed. This rerun uses the updated attribution code.

@cquil11
cquil11 force-pushed the codex/pr56318-bd57138-agentx branch from 593f473 to 1d542c7 Compare September 26, 2026 20:25
…cipe

Replace the hand-rolled bash server script
(dsv4_fp4_b200_vllm_cache_sources_mtp.sh) with a native srt-slurm
recipe at dsv4/vllm/b200-fp4-mtp/cache-sources.yaml. Three variants
select by CONC/KV_OFFLOADING: TP8 c8 native DRAM, TP8 c14 Simple NVMe,
TP8 c14 native DRAM+NVMe. NVMe variants use host_setup to
create/teardown /scratch/inferencex-kv-$SLURM_JOB_ID and
container_mounts with {job_id} templating.

Each search-space row in nvidia-master.yaml now carries srt-recipe:,
routing to the native-single-node launch path in the b200-nscale
launcher. The bash-specific routing, /ix mount override, NVMe directory
management, and GPU_MEMORY_UTILIZATION=0.85 export are removed from
launch_b200-nscale-slurm.sh (gpu-memory-utilization is now 0.85 in the
recipe).

A new srt-slurm patch (pr56318-overlay-validation.patch) adds
pr56318-overlay-check.sh to configs/patches/, referenced by the
recipe's setup_script to verify overlay checksums before engine start.

Co-Authored-By: Claude Opus 4.6 <[email protected]>
@github-actions

github-actions Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

@cquil11

cquil11 commented Sep 27, 2026

Copy link
Copy Markdown
Collaborator Author

/use 36272389390

@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

3 participants